Papers with speech technologies

4 papers
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for learning meaningful representations from unannotated data are resource-intensive and degrade other speech components.
Approach: They propose a method that decomposes SSL representations into speaker-specific components and generates speaker disentangled representations.
Outcome: The proposed method achieves speaker independence and improves on state-of-the-art methods.
Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond (2025.naacl-long)

Copied to clipboard

Challenge: Afrispeech-Dialog is a benchmark dataset of 50 simulated medical and non-medical African-accented English conversations . a 10%+ performance degradation is found in ASR systems on long-form, accented speech .
Approach: They propose to use a dataset to evaluate automatic speech recognition systems on African-accented conversations.
Outcome: The proposed dataset compares state-of-the-art speech recognition systems on accented conversations with native accents and shows a 10%+ performance degradation.
Open-source Multi-speaker Corpora of the English Accents in the British Isles (2020.lrec-1)

Copied to clipboard

Challenge: Using a dataset of high-quality audio, the authors examine the accents of 120 volunteers in the British Isles.
Approach: They present a dataset of high-quality audio of English sentences recorded by volunteers with different accents of the British Isles.
Outcome: The transcribed audio includes pronunciations of global locations, major airlines and common personal names in different accents.
Phonetic Segmentation of the UCLA Phonetics Lab Archive (2024.lrec-main)

Copied to clipboard

Challenge: ''big data'' does not exist for the majority of the world's languages . a corpus of audited phonetic transcriptions and phone-level alignments is available for free .
Approach: They present a corpus of audited phonetic transcriptions and phone-level alignments from the UCLA Phonetics Lab Archive . they discuss the utility of the corpus for general research and pedagogy in crosslinguistic phonetics .
Outcome: The VoxAngeles corpus improves the original corpus for phonetic typology and word- and phone duration measurements.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations